JMIR Public Health and Surveillance
◐ JMIR Publications Inc.
Preprints posted in the last 90 days, ranked by how well they match JMIR Public Health and Surveillance's content profile, based on 45 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Sadeghi Naieni Fard, F.; Oppong, J. R.; Tiwari, C.; Boakye, K.; Fard, F.
Show abstract
Cancer prevalence is distributed unevenly across regions and caused by the interaction of multiple risk factors. Previous studies focused on the use of global modeling techniques to predict cancer at the county level that overlooks important spatial differences. This study aims to develop geographically weighted machine learning models to predict cancer prevalence at the census tract level in the United States and identify local determinants of cancer burden. First, a scoping review was conducted to find a list of measurable drivers of cancer in the United States. Using this list, the data of these variables for 84415 census tracts were obtained from the Center for Disease Control and Prevention PLACES dataset and other publicly accessible resources. Then, several predictive models, including Ordinary Least Squares (OLS) and Geographically Weighted Regression (GWR), as well as Random Forest, XGBoost, and Deep Neural Network and their geographically weighted counterparts, were developed and compared using the Coefficient of Determination, Root Mean Square Error, and Absolute Error. Results presented that geographically weighted models outperformed other methods, and geographically weighted XGBoost achieved the strongest and most consistent overall performance with pseudo-R2 ranging between 0.89 and 0.98. Feature importance analysis of this model illustrated that most important cancer drivers changed location by location. Aged people, racial composition, preventative behaviors, and metabolic conditions such as diabetes, hypertension, and high cholesterol were determined as influential predictors, although their relative importance varied across regions. These findings revealed the value of localized models at a small geographic scale to identify regional cancer risk patterns and help the allocation of proper resources to hotspot areas. Keywords: Cancer prevalence, Census tracts, geographically weighted machine learning models, Deep neural network, XGBoost, Random Forest, Ordinary Least Squares, risk factor, determinant
Ohno, K.; Hirai, M.; hashimoto, s.
Show abstract
Background: In Japan, health planning is organized around secondary medical areas (SMAs; niji-iryo-ken; 330 areas in the 2025 classification), yet nationwide analyses of intensive care unit (ICU) capacity have been conducted mainly at the prefecture level, and a recent SMA-level study addressed only the presence or absence of ICUs. The full supply structure of intensive and intermediate critical care - ICU and high care unit (HCU) beds - has not been characterized at the SMA level with respect to its composition, road-network accessibility, and evolution over time. Methods: We developed MeshScope-Region, an analytical platform built on the Hospital Bed Function Reports (byosho-kino-hokoku) for fiscal years 2016-2024, in which ICU and HCU beds were identified from notified reimbursement categories and aggregated to SMAs. Three analytical layers were integrated: (1) cross-sectional distribution of ICU/HCU beds; (2) nationwide road-network accessibility computed with the Open Source Routing Machine (OSRM) from 176,962 populated 1-km census grid cells to all facilities reporting ICU or HCU beds; and (3) a nine-year longitudinal analysis of supply-structure types, classified by k-means (k = 6) in an 8-dimensional PCA space anchored to fiscal year 2024, with earlier years projected into the same space. Results: In fiscal year 2024, 20,631 ICU/HCU beds were reported nationally (7,114 ICU-type; 13,517 HCU-type) at 1,044 facilities. Zone-level totals among SMAs with any beds ranged 229-fold (3-688 beds); the 90th/10th percentile ratio of per-capita density was 3.6. In total, 90.1% of the population resided within 30 minutes' drive of a facility with ICU beds and 97.8% within 60 minutes; only 0.8% resided beyond 90 minutes. Although 140 of the 330 SMAs had no ICU facility within their own boundaries, 84.7% of their residents could reach an ICU facility in an adjacent area within 60 minutes' drive. Longitudinally, supply structures were highly persistent: 63.0% of SMAs (208/330) retained the same structural type across all nine years, adjacent-year rank correlations of a supply-vulnerability index were 0.887-0.924 (2016 vs. 2024: rho = 0.711), and the number of SMAs with zero ICU beds remained frozen at 133-141. The Gini coefficient of bed distribution declined from 0.384 to 0.262 - although computed on ICU-type beds alone it remained 0.365 in fiscal year 2024 - and capacity growth (total +27.9%) was driven predominantly by HCU beds (+41.6%) while ICU beds grew only +8.0%. Conclusions: Japan's critical care supply structure is regionally rigid, with a stable set of approximately 140 SMAs lacking ICU beds for nearly a decade, yet road-network accessibility substantially mitigates the consequences of zone-level absence. Recent capacity growth - and much of the apparent equalization - has occurred predominantly in intermediate care. MeshScope-Region provides a standing, reproducible evidence base at the geographic unit of Japan's medical planning cycles.
Corona-Moreno, R.; Acuna-Zegarra, M. A.; Santana-Cibrian, M.; Velasco-Hernandez, J. X.
Show abstract
During the COVID-19 pandemic, limited testing capacity and reporting delays complicated epidemic surveillance and decision-making in Mexico. We calibrated \textit{covidestim}, a Bayesian nowcasting model, to estimate the total SARS-CoV-2 infections from reported cases and deaths using Mexican surveillance data. Disease-progression distribution priors were calibrated using Mexico City records and validated through comparisons with national seroprevalence surveys, hospitalization data, and annual reported severe-case rates across all states. Using the reconstructed estimates of active infections, we implemented an event-based risk framework that quantifies the probability of encountering at least one infectious individual in gatherings of different sizes. This probability was subsequently translated into a four-level epidemiological traffic-light indicator and computed at both state and municipality levels. The resulting estimates revealed substantial spatial heterogeneity that is obscured by state-level aggregation, particularly in states with marked differences between urban and rural municipalities. To evaluate consistency with public-health indicators, we compared the proposed risk classification with the official Mexican epidemiological traffic-light system, considering interpretable gathering sizes relevant to public-health decision making. Weekly reports derived from this framework were delivered to policymakers in the State of Queretaro in Mexico, as an anticipation tool for school reopening and public-space management. This demonstrates that this Bayesian reconstruction of infections combined with event-based risk metrics can provide an interpretable and generalizable municipality-level complement to routine surveillance systems, particularly in regions with limited testing capacity and heterogeneous local transmission dynamics.
Yahaya, Y.; Khan, S.; Rani Saha, P.; Meia, M. A. A.
Show abstract
Diagnosed diabetes affects approximately 38.4 million Americans, but its burden is not evenly distributed across U.S. counties. Existing machine-learning studies have mainly focused on individual risk prediction using biometric, clinical, or survey variables. These approaches are less suited to explaining why diagnosed diabetes prevalence differs geographically across counties. We developed an explainable gradient-boosting framework for predicting county-level diagnosed diabetes prevalence across 2,957 U.S. counties using an ecological cross-sectional design. The analysis integrated food-environment, socioeconomic, occupational, demographic, health-behavior, and clinical indicators from five public data sources. Four regression models were compared: Elastic Net, Random Forest, XGBoost, and LightGBM. LightGBM was selected as the primary model based on validation-set RMSE and interpreted using SHAP TreeExplainer. The validation-selected LightGBM model achieved a held-out test RMSE of 0.423 percentage points, R{superscript 2} = 0.964, and MAPE = 2.76%. Although XGBoost achieved a lower test RMSE of 0.399 and R{superscript 2} = 0.968, it was retained as a secondary benchmark because primary-model selection was based only on validation performance. A sensitivity model using only structural and contextual predictors, and excluding CDC PLACES health-behavior and clinical covariates, retained substantial predictive performance (R{superscript 2} = 0.827). Poverty rate was the most frequent dominant positive structural SHAP contributor nationally (n = 772 counties, 26.1%), followed by food insecurity rate (n = 707, 23.9%), Supplemental Nutrition Assistance Program (SNAP) participation rate (n = 316, 10.7%), unemployment rate (n = 224, 7.6%), and median household income (n = 178, 6.0%). Residual Morans I decreased from 0.665 to 0.069 after model fitting. Explainable machine learning using public county-level data can characterize geographic variation in diagnosed diabetes prevalence. County-level SHAP maps may support local hypothesis generation, but should be interpreted as explanations of model predictions rather than causal effects.
Gao, J.; Windett, J. H.; Ademu, L. O.; Li, Z.; Idris, M. A.; Griffin, B. C.; Zhang, Y.; Radford, B. J.
Show abstract
Background Human papillomavirus (HPV) vaccination is an effective cancer prevention strategy, yet HPV vaccine awareness remains uneven across sociodemographic groups. In the current digital information environment, awareness may be shaped not only by access to health information but also by exposure to false or misleading health information, difficulty evaluating information accuracy, and echo-chamber dynamics on social media. Objective This study examined associations between perceived exposure to false or misleading health information on social media, difficulty determining whether social media health information is true or false, perceived echo-chamber exposure, and HPV vaccine awareness among U.S. adults. Methods We analyzed nationally representative Health Information National Trends Survey data using survey-weighted descriptive statistics and logistic regression models. The analytic sample included 2,371 respondents, representing a weighted population of 49.2 million U.S. adults. The outcome was HPV vaccine awareness. Primary predictors included perceived exposure to false or misleading health information on social media, difficulty determining whether social media health information was true or false, and perceived same-view health network exposure on social media. Models adjusted for age, sex, race/ethnicity, education, household income, rurality, and Census division. Results Overall, 60.37% of respondents reported HPV vaccine awareness. Most respondents reported encountering false or misleading health information on social media, with 45.57% reporting "some" and 32.83% reporting "a lot." In unadjusted models, greater perceived exposure to false or misleading health information was associated with higher odds of HPV vaccine awareness. After adjustment, respondents reporting "some" false or misleading health information had significantly higher odds of HPV vaccine awareness compared with those reporting none (AOR=2.40, 95% CI: 1.06-5.43), while the association for "a lot" was marginal (AOR=2.29, 95% CI: 0.97-5.38). Difficulty identifying true versus false social media health information and perceived echo-chamber exposure were associated with HPV vaccine awareness in unadjusted models but were attenuated after adjustment. HPV vaccine awareness was substantially higher among females and respondents with higher educational attainment, and lower among Hispanic, non-Hispanic Asian, and non-Hispanic other respondents compared with non-Hispanic White respondents. Conclusions HPV vaccine awareness is associated with both digital health information exposure and persistent sociodemographic inequities. Greater perceived exposure to misleading health information may reflect broader engagement with health-related content on social media, where accurate and inaccurate information coexist. Public health communication strategies should address misinformation vulnerability while expanding accurate, culturally responsive HPV vaccine messaging across digital platforms.
Wang, K.; Olaniyan, P.; Powla, P.; Pabon-Rodriguez, F. M.
Show abstract
Indiana still faces significant health challenges, ranking among the least healthy U.S. states due to high obesity rates, mental health issues, and other chronic conditions. These disparities are closely linked to inequities in healthcare access, which are largely shaped by social determinants of health. Using data from the Social Vulnerability Index and County Health Rankings and Roadmaps, this study analyzes trends in obesity, mental health, and premature death across Indiana counties before, during, and after the COVID-19 pandemic. Descriptive statistics, correlation analyses, and Negative Binomial regression models were used to evaluate county-level disparities. In 2018, higher rates of uninsured, obese, and physically inactive populations were associated with increased premature death. In 2020, diabetes, smoking, and alcohol consumption were significant factors. By 2022, unemployment, education, obesity, insurance, exercise access, and mental health provider availability were associated with premature death. Findings indicate that socially vulnerable counties experienced amplified health impacts, with obesity rising most sharply where exercise infrastructure was limited and poor mental health days increasing across all counties. These results highlight persistent service gaps and the critical need for targeted investments in recreational infrastructure and mental healthcare. Future research should examine policy influences and causal relationships to inform equity-focused interventions.
Kathuria, Y.; Miller, K.; Selden, E. B.; Gallagher, W. J.; Capan, M.
Show abstract
Patients diagnosed with type 2 diabetes (T2D) are at increased risk of developing cardiovascular disease (CVD), the leading cause of morbidity and mortality in this population. Early detection and glycemic control within the first year after diagnosis reduce CVD risk. However, gaps remain in how to operationalize early detection of T2D using Electronic Health Record (EHR) data and quantify its relationship with subsequent CVD risk using longitudinal observations. We developed a probabilistic graph model to analyze the interdependencies between early detection of T2D, post-diagnosis glycemic control, and CVD occurrence. Using a temporally structured Bayesian Network (BN) learned from EHR data of 9,450 primary care patients between 2017 and 2023, we quantified probabilistic dependencies between demographics, diagnostic delay surrogates, glycemic control, and post-diagnosis CVD occurrence. Percentile based thresholds defined risk groups, where individuals with predicted probabilities in the bottom decile ([≤] 10th percentile) were classified as low risk, and those in the top decile ([≥] 90th percentile) as high risk. Results demonstrated heterogeneity in predicted risks across glycemic and cardiovascular outcomes. Predicted probability of developing CVD within the first year after T2D diagnosis ranged from a mean of 5.2% in the low-risk group to 28.9% in the high-risk group, while predicted probabilities of mean Hemoglobin A1c (HbA1c) [≥] 8% during the first year post-diagnosis ranged from 1.6% in low-risk to 55.1% in high-risk group. Patients with HbA1c at diagnosis [≥] 8% had higher predicted probabilities of first-year post-diagnosis mean HbA1c [≥] 8% (53.3% vs. 1.9%) and high HbA1c coefficient of variation (18.7% vs. 3.1%) compared with those with HbA1c [≤] 6.5%. Incorporating early clinical outcomes refined later risk predictions, with long-term CVD risk reaching 33.5% among high-risk individuals. The proposed model achieved predictive performance comparable to conventional machine learning approaches while providing interpretable relationships for risk stratification in primary care populations.
Muharram, F. R.; Zulfikar, M. Q. B.; Siregar, R. A.; Nur, A.; Widyahening, I. S.; Danaei, G.
Show abstract
ABSTRACT Background: To examine trends in Indonesia's diabetes care cascade from 2013 to 2023, identify key determinants, and assess progress toward global targets of 80% diagnosis and 80% glycemic control among those diagnosed. Methods: We analyzed nationally representative data from Indonesia's Health Surveys in 2013, 2018, and 2023. Diabetes was defined using fasting plasma glucose and oral glucose tolerance tests. We estimated diagnosis, treatment, and control rates and examined sociodemographic predictors of cascade progression using survey-weighted logistic regression models. Results: Between 2013 and 2023, the prevalence of diabetes among adults aged [≥]15 years remained stable, ranging from 10.7% to 11.8%. Diagnosis increased from 15.1% (95% CI: 13.4-16.7) to 20.7% (18.5-22.9), treatment nearly doubled from 10.5% (9.1-11.9) to 19.0% (16.9-21.2), and control rose modestly from 4.6% (3.6-5.6) to 6.5% (5.2-7.8). Older age, urban residence, higher socioeconomic status, and insurance coverage were associated with greater progression through the cascade. Wealth-related inequalities persisted in 2023: one-third of cases were diagnosed among the richest (35.3% [29.0-41.7]) versus only 11.0% (7.9-14.2%) among the poorest. Compared with the lowest quintile, wealthier individuals had higher odds of diagnosis (AOR 3.55 [2.11-5.98] for diagnosis, 3.59 [2.01-6.41] for treatment, and 2.24 [1.12-4.51] for control). Conclusions: Indonesia achieved meaningful improvements in the diabetes care cascade over the past decade, yet remains far below global 80/80 targets, with nearly 80% of cases undiagnosed and control below 10%. Persistent wealth-based inequities highlight that near-universal insurance coverage has not been translated into equitable care access, underscoring the need for equity-focused screening and primary care strengthening. Keywords: Diabetes, Care Cascade, Health Services, Indonesia
Jaganath, D.; Ilavarasan, V.; Wong, R.; Chitnis, A.; Murrill, M. T.
Show abstract
Context: Most individuals in the United States have commercial health insurance, yet costs for tuberculosis (TB) care have focused on the public sector. Objective: To quantify 12 month all cause healthcare costs and identify predictors of expenditure among commercially insured persons with TB disease in the United States. Design/Setting: Retrospective cohort study using Merative (TM) MarketScan (R) Commercial Claims Database (2013 to 2018). Participants: Adults 18 years old with TB disease Main Outcome Measure: Total 12 month all cause healthcare costs (outpatient, inpatient, pharmacy) were calculated from the date of diagnosis. Adjusted cost ratios (aCR) were estimated using a Gamma generalized linear model. Results: We included 303 individuals diagnosed with TB disease, median age 46 years, 158 (52%) male, 16 (5%) with HIV, 12 (4%) with hepatitis B (HBV), and 13 (4%) with a drug use disorder. Mean total 12-month costs were $32,404 (median $8,075; SD $78,829). Median 12-month costs were substantially higher among persons with any comorbidity (HIV, HBV, hepatitis C (HCV), alcohol use disorder, drug use disorder, or Charlson score >0) compared to those without ($11,930 [IQR $4,194 to $36,073] vs $3,385 [IQR $1,506 to $8,609]; p<0.001). HIV coinfection and drug use disorder were the strongest independent predictors. HIV coinfection was associated with 4.7 fold higher costs (aCR 4.70, p<.001), driven predominantly by pharmacy expenditure (aCR 16.4). Drug use disorder was associated with 3.2 fold higher costs (aCR 2.62, p=.03). Comorbidity burden was a continuous independent predictor (aCR 1.36 per Charlson point, p<.001). Conclusions: Healthcare costs are high among persons with TB who have commercial insurance, and are further increased with comorbidities including HIV coinfection and drug use disorder. Improved screening, care coordination and management of TB and high risk comorbidities could yield significant cost savings.
Dojcsak, L.; Abegaz, T.; Islam, M.; Chandler, Y.; Maleku, A.; Doubeni, A.; Mohammed, B.; Langston, M. A.; Donneyong, M. M.
Show abstract
Health-related social needs (HRSNs), such as housing instability, food insecurity, and transportation challenges, are nonmedical factors associated with poorer health and well-being. Screening for unmet HRSNs is a critical step towards identifying at-risk patients, but manual screening is resource intensive and often incomplete. We utilized Electronic Health Records (EHR) data to develop machine learning models to identify unmet HRSNs using a limited set of non-modifiable sociodemographic features available in EHRs. We included 745,975 patients screened for at least one HRSN using data from community health centers that participated in the OCHIN practice-based research network between 2016 and 2022. Logistic regression, random forest (RF), eXtreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM) algorithms were trained to predict unmet HRSNs. Model performance was evaluated using 10-fold cross-validation and area under the receiver operating characteristic curve (AUROC). For overall HRSN prediction, LightGBM (AUROC, 64.5%, 95%CI: 64.3, 64.7) performed slightly better than logistic regression (61.4%), RF (63.7%), and XGBoost (60.3%). Similar performances were observed predicting individual HRSNs. Model performances were modest; however, they establish a benchmark for predictive performance achievable using only routinely available demographic data and provide a foundation for incorporating additional clinical and area-level social determinants of health data.
Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.
Show abstract
Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.
Hwang, Y.-M.; Cui, Y.; Xu, J.; Pan, T.; Li, R.; Rice, B.; Tian, L.; Hernandez-Boussard, T.
Show abstract
Comorbidity indices are widely used in clinical research to summarize disease burden. However, traditional indices were developed decades ago in limited populations using fixed weights that do not reflect the diversity of patients in modern healthcare. We present the Personalized Comorbidity Score (PCS), a data-driven framework for context-dependent comorbidity scoring designed to capture patient complexity while remaining accessible for broad research adoption. PCS was developed using Epic Cosmos, a large national EHR network encompassing over 8 million adult inpatient encounters from 2015 to 2020, with comorbidities defined using AHRQ Clinical Classifications Software Refined categories. Models were developed separately across eight age-sex subgroups using LASSO-penalized Cox regression for feature selection and restricted mean survival time for score derivation. PCS is available in two versions: PCS Core, incorporating age, sex, and comorbidities, and PCS Extended, which additionally incorporates socioeconomic and geographic variables. PCS Core and PCS Extended achieved AUROCs of 0.812 and 0.813 for one-year mortality, outperforming traditional indices (AUROC, 0.714-0.730). PCS demonstrated consistently lower subgroup calibration error across demographic and socioeconomic groups without including race or ethnicity as model features. PCS was further evaluated in two complementary external EHR datasets (Stanford Health Care and MIMIC-IV) with distinct patient populations and data structures, where it consistently outperformed traditional indices. Open-source R and Python packages are provided to support broad adoption. PCS provides an updatable framework for comorbidity measurement that is accurate, context-dependent, and designed to evolve alongside clinical practice.
Srivastava, D. K.; Gupta, S.; Yadav, N.
Show abstract
Background: Evaluation of public health surveillance systems is a programmatic obligation but has largely been conducted as a periodic, externally commissioned activity requiring dedicated resources and additional data collection. India's Integrated Disease Surveillance Program (IDSP) generates continuous outbreak data through weekly reports but lacks a routine, embedded performance evaluation mechanism. This study assessed the quality of IDSP outbreak detection and response across multiple surveillance attributes and developed a weighted composite performance scoring framework using only routine program data. Methods: A cross-sectional evaluation study was conducted across 38 districts of Bihar using secondary data from IDSP Central Surveillance Unit weekly outbreak reports for 2016 - 2018 (n=559 outbreaks). Six surveillance quality attributes were assessed - timeliness, completeness, representativeness, relative sensitivity, acceptability and flexibility. A weighted composite performance scoring scale was developed using expert opinion-derived attribute weightages (n=25 experts). District-level scores were computed and scaled to 100. Results: Timeliness was the poorest-performing attribute, with fewer than 15% of outbreaks notified within 48 hours across all three years. Private sector participation was entirely absent - the acceptability score was 0 across all 38 districts for all three years. Completeness was the strongest attribute, exceeding 95% in all years. The mean composite score remained consistently low (23 - 27 out of 100) with widening inter-district disparity over time. Four districts (10.5%) scored 0 in all three years. Conclusions: This study presents a dynamic, routine-data-based composite performance evaluation framework for IDSP outbreak detection and response. The modular, configurable framework functions at any administrative level (from block to national) and is compatible with digital health information platforms, enabling continuous, embedded performance monitoring without additional data collection. The framework has been registered as an Intellectual Property with the Government of India. Keywords: Disease surveillance; IDSP; IDSR; performance evaluation; composite score; outbreak detection; timeliness; completeness; relative sensitivity; digital health
Garavito Jimenez, D. A.; Bello Angulo, D. E.; Mejia Lemus, L. T.; Chipatecua, D.; Fula, D. D.; Perez-Rubiano, S.; Martinez, F. L.; Bohorquez Pinzon, J. C.
Show abstract
Between 2024 and 2025, Colombia universalized the Electronic Health Invoice with embedded Individual Health Services Delivery Records (RIPS -- Registro Background Between 2024 and 2025, Colombia universalized the Electronic Health Invoice with embedded RIPS records (FEV-RIPS) as the standard for financial and clinical data exchange. ADRES -- the entity responsible for administering the resources of Colombia's General Social Security Health System -- faced the challenge of processing information from multiple heterogeneous sources generated by more than 55,000 healthcare providers. Health systems in high-income countries converge clinical-financial data in consolidated platforms; Colombia started from a fragmented architecture with incompatible historical sources, no cross-database standardization, and no centralized analytical infrastructure until 2023. Objective We describe the design, technical challenges of integrating heterogeneous data, and operational performance of the analytical infrastructure built by ADRES to centralize large-scale processing of Colombian health system information, and derive transferable lessons for health system resource administrators in Latin America facing equivalent digitalization mandates. Methods Technical-descriptive report based on operational metrics from the ADRES Azure/Databricks environment during January-November 2025. We report indicators of data volume, processing speed, computational capacity, concurrent use by functional group, and governance structure. The architecture integrates VPN connectivity with MinSalud, automated processing of multiple formats (XML, relational tables, flat files), and a medallion data lake (Bronze/Silver/Gold). Data quality challenges include structural inconsistencies across sources, coding incompatibilities (municipalities, dates, diagnoses), format heterogeneities in unstructured data, and absent technical documentation. Results The platform manages 21 catalogs, 1,183 tables, and over 110,645 million stored records, with cumulative production exceeding 1 trillion processed records. It executes queries on 100 billion records in ten seconds using clusters of up to 32 TB RAM and 4,096 vCPU. During September-October 2025, monthly query peaks reached 78,028 across eleven functional groups. Integration required Python/PySpark parsers for variable-depth XML, equivalence tables for incompatible municipality codes, cleaning routines for extreme dates used as nulls (1900-01-01, 9999-12-31), and transformation logic bridging classic RIPS and FEV-RIPS. The platform supported econometric analyses, judicial mandate responses, and public interactive dashboards. Conversational AI integration (Genie, Copilot) extends analytical access to users without SQL knowledge. Conclusions ADRES built in one year an analytical infrastructure that provides, to our knowledge, the first published documentation of the systemic technical challenges of integrating heterogeneous data sources in a middle-income social security health system. Centralizing health system information at national scale is technically feasible under public institutional constraints -- but requires solving cross-source standardization problems the implementation literature does not document with quantitative precision. The derived lessons are transferable to health system resource administrators in Latin America facing equivalent challenges.
Zhou, Y.; Ma, J.; Zhang, Y.
Show abstract
Endometrial cancer (EC) incidence is closely linked to metabolic and hormonal factors. The TyGFI, a composite indicator integrating the triglyceride-glucose index and frailty index, may capture combined risk dimensions relevant to EC etiology and prediction. This study aimed to evaluate the association between TyGFI and EC prevalence among U.S. women aged 45 years and older, and to explore its predictive utility using machine-learning approaches. Data were drawn from the National Health and Nutrition Examination Survey 2011-2018 cycles. The exposure was TyGFI, and the outcome was EC status ascertained from self-reported cancer history and standardized questionnaires. From an initial 39,156 participants, we excluded males, individuals aged under 45 years, those missing TyGFI components or EC data, and extreme TyGFI values, yielding a final cohort of 2,837 women. We performed exploratory feature selection and built six predictive models using machine-learning algorithms to identify key predictors and evaluate TyGFI's contribution to model performance. Survey-weighted multivariable logistic regression estimated the association between TyGFI and EC prevalence with sequential adjustment for covariates. In a weighted sample representing 30,489,082 U.S. women aged 45 years and older, higher TyGFI was significantly associated with EC prevalence. After full adjustment, each unit increase in TyGFI corresponded to a 57% increase in the odds of EC (odds ratio 1.570; 95% confidence interval 1.033-2.370; p = 0.0322), with a significant dose-response trend across quartiles (p for trend = 0.0257). Among six machine-learning models, CatBoost achieved the highest predictive performance, with an area under the curve of 0.999. SHAP analysis identified TyGFI as the most influential predictor, followed by age and serum albumin. In this nationally representative sample of middle-aged and older U.S. women, TyGFI was significantly associated with EC prevalence and emerged as the dominant predictor in machine-learning models. These findings suggest that TyGFI may enhance risk stratification for EC beyond established reproductive and metabolic factors, though prospective studies are warranted to validate its clinical utility.
Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;
Show abstract
Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.
Braga, J. S.; Coelho, F. C.; Laiate, B.
Show abstract
Public health professionals have access to more data than ever before. Yet answering a relatively simple epidemiological question often requires navigating multiple databases, formats, software tools, and reporting systems. As a result, valuable data often remain locked behind technical barriers, making it harder for public health professionals to turn information into decisions. We developed EpidBot to simplify this process. EpidBot is a platform that allows users to retrieve, analyze, visualize, model, and report epidemiological data through natural language interaction. By connecting multiple public health data sources within a single environment, the platform enables users to conduct analyses that would traditionally require several independent tools and specialized technical skills. Rather than functioning solely as a search interface, EpidBot supports complete analytical workflows. Users can explore surveillance data, compare trends across locations and time periods, generate maps and visualizations, construct epidemiological models, and produce structured technical reports while maintaining full visibility of data sources and analytical procedures. To show what this looks like in practice, we present representative use cases, including the automatic generation of a mathematical model for Ebola virus disease in the Democratic Republic of the Congo. From a single user request, EpidBot assembled evidence from published sources, generated and calibrated a compartmental transmission model, identified key transmission drivers, evaluated intervention scenarios, and produced a technical report with quantitative findings and policy-relevant recommendations. EpidBot shows how natural language interaction can reduce the technical barriers that often separate public health professionals from the analyses they need to perform. By bringing data access, analysis, modeling, visualization, and reporting into a single environment, the platform helps transform information into evidence while preserving transparency and reproducibility.
Imahashi, M.; Noda, T.; Omata, K.; Yokomaku, Y.; Taniguchi, T.
Show abstract
Objective: In Japan, antiretroviral therapy (ART) for individuals living with human immunodeficiency virus (HIV) is financially supported through the Physical Disability Certification System for Immunological Impairment. However, certification requires multiple laboratory assessments after diagnosis, possibly delaying ART initiation. This study examined the impact of these eligibility requirements on ART initiation using real-world clinical data. Design: Single-center retrospective cohort study. Setting: Nagoya Medical Center, Japan. Subjects, participants: A total of 568 patients who attended their first consultation between 2015 and 2019 were included. Of these, 434 were ART-naive, and 134 had already initiated ART at the first visit. Main outcome measures: ART initiation rate, time to treatment initiation, factors associated with treatment delay, and utilization of the Physical Disability Certificate system. Results: Among the 434 untreated patients, the median time to ART initiation was 42 days. Seven patients (1.6%) did not meet the Grade 4 certification requirements and remained untreated. Overall, 13 of the 568 patients (2.3%) were affected by the certification system, including those importing ART from overseas or using alternative financial support mechanisms. Non-Japanese nationality, lack of health insurance, unstable employment, and low CD4 cell counts were significantly associated with failure to initiate treatment. Among the 134 previously treated patients, 108 (80.5%) had obtained a Physical Disability Certificate. Conclusions: Although relatively few patients were affected, certification requirements may delay ART initiation among socioeconomically vulnerable populations. Further multi-center and cost-effectiveness studies are needed to improve compatibility between long-term financial support systems and rapid ART initiation strategies after diagnosis.
Chicoine, G.; Germain, N.; Turcotte, S.; Cote, E.; Gelinas, V.; Legare, F.; Paquette, J.-S.; Totten, A. M.; Morin, M.; Straus, S. E.; Archambault, P. M.
Show abstract
Purpose: Serious Illness Conversations (SICs) are essential to delivering person-centered care for older adults with chronic conditions, but are rarely integrated into routine primary care. To address this gap, we compared the effectiveness of a structured training strategy versus passive dissemination of educational materials on SIC documentation rates during the COVID-19 pandemic. Methods: A quasi-experimental study across 13 primary care clinics in Quebec, Canada. Five clinics received structured team-based Serious Illness Care Program training (intervention group) with a provincially disseminated SIC toolkit and eight received the toolkit only (control group). The primary outcome was the proportion of patients with a documented SIC across three time periods (Period 1, pre pandemic; Period 2, pandemic initial wave; and Period 3, post dissemination of SIC toolkit). We used generalized estimating equations (GEE). Results: Across 13 clinics, 2,368 eligible patients (mean age 75.8 years (SD = 7.5), 54% female, with a mean Charlson Comorbidity Index of 4.88 (SD = 2)) accounted for 19,134 clinical visits, 49.5% in person and 49.6% virtually. SIC documentation rates were 3.3% (control) and 3.4% (intervention) in Period 1, 9.3% and 4.3% in Period 2, and 6.4% and 4.8% in Period 3, respectively. There was no statistically significant improvement to SIC documentation in the intervention group at Period 2 nor Period 3. Conclusion: Structured training was not more effective than passive dissemination for SIC documentation. Educational interventions must be supported by structural changes, workflow integration, and organizational leadership. Multi-level implementation strategies are needed to embed SICs sustainably into primary care.
Hilliard, M. E.; Foreman, R.; Khan, T.; Zona, E.; Mishra, A.; Howse, S. J.
Show abstract
Background: For US young adults aged 18-25 in the 2018-2024 period, fentanyl was involved in 78.2% of the 44,020 unintentional or undetermined-intent overdose deaths, most often co-involving stimulants and other non-opioid substances. While fatal overdose rates in this age group have fallen to their lowest recorded level, emergency medical services-attended non-fatal overdose events have reached record highs, shifting the decisive variable toward bystander recognition and response. College students report near-universal alcohol education but minimal education on the substances actually driving overdose mortality. Methods: We conducted a single-group pre-post evaluation of the DopaGE Portal, a gamified, mastery-based digital platform covering cocaine, MDMA, benzodiazepines, and opioid overdose response, deployed at a public university (UNL) and a multi-campus volunteer network (TACO). Paired pre/post surveys (N=42) measured self-efficacy (7 items; primary), behavioral intentions, risk perception, and knowledge/attitudes on 5-point scales, plus four factual knowledge questions. Paired t-tests, exact McNemar tests, and Benjamini-Hochberg correction across eight primary tests were applied. Institutional naloxone distribution at UNL was tracked as an ecological behavioral outcome. A mandated high-school cohort (N=94) provided supplementary acceptability data. Results: Self-efficacy increased from 2.82 to 4.46 (d=2.00, 95% CI 1.46-2.55; adjusted p<.001), and behavioral intentions from 4.24 to 4.81 (d=1.43; adjusted p<.001), with effects statistically indistinguishable across sites. Three of four knowledge items improved significantly (+31 to +41 percentage points). Risk perception was at ceiling at baseline (4.38/5) and did not change. In the two months following deployment, 38 naloxone kits were distributed on campus (limited to one per person from the campus pharmacy and health center) versus 14 in the preceding two years combined; the campus health center had distributed zero kits in 2025 despite stocked availability. Evaluation ratings were uniformly positive across voluntary and mandated cohorts, with zero negative ratings. Conclusions: A digital-only, gamified intervention produced large gains in overdose-response self-efficacy and substance-specific knowledge, with concurrent campus-level naloxone acquisition consistent with behavioral translation. These findings are preliminary -- single-group, modest N, ecological behavioral outcome -- and motivate a future randomized controlled trial.